Skip to content

fix(codex): keep system prompts in input for GPT-5 automatic prompt caching - #1346

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.6.7from
Gi99lin:fix/codex-prompt-caching
Apr 16, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.6.7from
Gi99lin:fix/codex-prompt-caching

Conversation

@Gi99lin

@Gi99lin Gi99lin commented Apr 16, 2026

Copy link
Copy Markdown
Contributor

Problem

OpenAI's automatic prompt caching for GPT-5 models returns cached_tokens: 0 on every request routed through OmniRoute, even with identical 53K-token prompts sent seconds apart.

Root cause: The instructions field in the Responses API is not included in the prompt cache key computation for GPT-5 models. OpenAI only caches based on the serialized input array + tools. The current code moves system messages from input (cacheable) into instructions (not cacheable) via hoistSystemMessagesToInstructions(), destroying the cache prefix entirely.

Evidence

Diagnostic logging on a production deployment confirmed:

[cache-diag-resp] model=gpt-5.4 input=52889 output=370 cached=0 pct=0%   (request 1)
[cache-diag-resp] model=gpt-5.4 input=53421 output=29  cached=0 pct=0%   (request 2, 15s later)
[cache-diag-resp] model=gpt-5.4 input=53973 output=263 cached=0 pct=0%   (request 3, 50s later)

Three consecutive requests with identical instructions + tools prefix, same model, within 90 seconds — zero cache hits. Meanwhile, a MiniMax provider using the same proxy achieved cache hits immediately:

MiniMax request 1: cache_creation=10513, cache_read=0
MiniMax request 2: cache_creation=438,   cache_read=10513  ← cache hit!

This is consistent with community reports:

What this PR changes

For native Responses API passthrough requests (_nativeCodexPassthrough === true):

Aspect Before After
System messages Hoisted from input to instructions Converted system → developer role, kept in input
instructions field CODEX_DEFAULT_INSTRUCTIONS (3000+ words) Minimal placeholder
Cacheable prefix tools only (system prompt removed from input) developer msg + tools + conversation prefix
Expected cache rate 0% ~90%+ (matching direct connections)

For translated requests (Chat Completions → Responses format): no change. The existing hoist + default instructions behavior is preserved.

New function: convertSystemToDeveloperRole()

Converts role: "system" → role: "developer" in-place within the input array. This is necessary because:

  1. Codex rejects system role in input (existing comment in code confirms this)
  2. GPT-5 models support developer role as the replacement
  3. Keeping the content in input preserves the cacheable prefix

Impact

For a deployment routing ~7K daily Codex requests averaging 41K input tokens each:

  • Before: 286.7M input tokens/day, 0.1% cache rate, ~8 cache hits total
  • After (projected): Same volume, ~90% cache rate → ~258M cached tokens/day
  • Cost reduction: ~90% on input token costs for cached portions
  • Latency reduction: Up to 80% TTFT for cache hits (per OpenAI docs)

Testing

To verify the fix:

  1. Deploy the updated executor
  2. Monitor cached_tokens in response.completed SSE events
  3. Second request with the same system prompt should show cached_tokens > 0

Example diagnostic (add to proxy/middleware):

// In response handler, look for response.completed events:
const u = completed.response?.usage;
const cached = u?.input_tokens_details?.cached_tokens ?? 0;
console.log(`cached=${cached} pct=${Math.round(cached / u.input_tokens * 100)}%`);

OpenAI's automatic prompt caching for GPT-5 models computes the cache
key from the serialized `input` array and `tools`, but does NOT include
the `instructions` field. The current code moves system messages from
`input` into `instructions` via `hoistSystemMessagesToInstructions()`,
which removes them from the cacheable prefix entirely.

This causes 0% cache hit rates for native Responses API passthrough
clients that send system prompts as role=system messages in `input`.

Changes:
- Add `convertSystemToDeveloperRole()` — converts system → developer
  role in-place within `input` (Codex accepts developer but rejects
  system). This keeps the content in the cacheable prefix.
- For native passthrough requests: use the new function instead of
  hoisting to `instructions`. Set a minimal placeholder instruction
  instead of injecting CODEX_DEFAULT_INSTRUCTIONS.
- For translated requests (Chat Completions → Responses): preserve
  the existing hoist + default instructions behavior (no change).

Before: instructions="3000-word default + system prompt", input=[user msgs only]
  → cached_tokens = 0 (instructions not in cache key)

After:  instructions="minimal", input=[{role:"developer",...}, user msgs...]
  → cached_tokens = system_prompt + tools + conversation prefix

Ref: https://community.openai.com/t/caching-is-borked-for-gpt-5-models/1359574
Ref: https://community.openai.com/t/no-caching-with-model-responses/1338627
@Gi99lin
Gi99lin requested a review from diegosouzapw as a code owner April 16, 2026 18:31

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request implements a cache-aware strategy for system prompt handling in the CodexExecutor to optimize for GPT-5 models. It introduces the convertSystemToDeveloperRole function, which converts system roles to developer roles within the input array during native passthrough to preserve OpenAI's prompt caching. For translated requests, the existing behavior of hoisting system messages to the instructions field is maintained. I have no feedback to provide.

@diegosouzapw
diegosouzapw changed the base branch from main to release/v3.6.7 April 16, 2026 19:03
@diegosouzapw
diegosouzapw merged commit 79c63d1 into diegosouzapw:release/v3.6.7 Apr 16, 2026
2 checks passed
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @Gi99lin for this great contribution! 🎉 This PR has been evaluated and successfully integrated into the release/v3.6.7 branch and will be part of the final production release. We appreciate your effort!

@diegosouzapw diegosouzapw mentioned this pull request Apr 22, 2026
@Gi99lin
Gi99lin deleted the fix/codex-prompt-caching branch May 21, 2026 21:34
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants